Papers with natural language processing methods
T-Know: a Knowledge Graph-based Question Answering and Infor-mation Retrieval System for Traditional Chinese Medicine (C18-2)
Copied to clipboard
| Challenge: | Traditional Chinese Medicine (TCM) is one of precious intangible cultural heritages of the Chinese nation. |
| Approach: | They propose to use authorized and anonymized clinical records, medicine clinical guidelines, teaching materials, classic medical books, academic publications, etc. as data resources to build a TCM knowledge graph. |
| Outcome: | The proposed system extracts triples from free texts to build a TCM knowledge graph. |
EfficientOCR: An Extensible, Open-Source Package for Efficiently Digitizing World Knowledge (2023.emnlp-demo)
Copied to clipboard
| Challenge: | Existing OCR engines fail to provide accurate, cost-effective and sample-efficient character recognition for public domain documents. |
| Approach: | EffOCR is an open-source optical character recognition package that is accurate, cheap to deploy and sample efficient to customize to novel collections, languages, and character sets. |
| Outcome: | EffOCR model trains character retrieval problem and scales to novel collections, languages, and character sets. |
Casting Light on Invisible Cities: Computationally Engaging with Literary Criticism (N19-1)
Copied to clipboard
| Challenge: | Literary critics often attempt to uncover meaning in a single work of literature through careful reading and analysis. |
| Approach: | They propose to use a literary theory to analyze Italo Calvino's novel Invisible Cities to leverage contextualized representations to embed each city's description and use unsupervised methods to cluster embeddings. |
| Outcome: | The proposed method can be applied to Italo Calvino’s novel Invisible Cities . authors compare results to similarity judgments generated by human readers . |
Extracting Chemical-Protein Interactions via Calibrated Deep Neural Network and Self-training (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Several natural language processing methods have been used to extract interactions between chemicals and proteins from biomedical text data. |
| Approach: | They propose a method to extract chemical–protein interactions from biomedical text data . they use a pre-trained language-understanding model and calibration techniques to estimate uncertainty . |
| Outcome: | The proposed approach achieves state-of-the-art performance on the Biocreative VI ChemProt task while preserving higher calibration abilities. |
Diversity, Density, and Homogeneity: Quantitative Characteristic Metrics for Text Collections (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing descriptive statistics are inadequate to summarize text collections by quantitative measures. |
| Approach: | They propose a set of characteristic metrics that quantitatively measure the dispersion, sparsity, and uniformity of a text collection. |
| Outcome: | The proposed metrics are highly correlated with text classification performance of a renowned model, which could inspire future applications. |
Summarizing Patients’ Problems from Hospital Progress Notes Using Pre-trained Sequence-to-Sequence Models (2022.coling-1)
Copied to clipboard
| Challenge: | Problem list summarization requires a model to understand, abstract, and generate clinical documentation. |
| Approach: | They propose a task that summarises patients' main problems from daily progress notes using input from the provider's progress notes during hospitalization. |
| Outcome: | The proposed model outperforms two state-of-the-art seq2seq transformer architectures in summarizing patients' main problems from daily progress notes in the medical information mart for Intensive Care (MIMIC)-III. |
Towards Debiasing Sentence Representations (2020.acl-main)
Copied to clipboard
Paul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim, Ruslan Salakhutdinov, Louis-Philippe Morency
| Challenge: | Recent work has shown word-level embeddings reflect and propagate social biases present in training corpora. |
| Approach: | They propose a method to debias word embeddings to reduce biases at sentence level . they hope their work will inspire future research on characterizing and removing biase . |
| Outcome: | The proposed method reduces biases and preserves performance on downstream tasks such as sentiment analysis and natural language understanding. |
Constructing Word-Context-Coupled Space Aligned with Associative Knowledge Relations for Interpretable Language Modeling (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to train language models have limitations in interpretability . a Word-Context-Coupled Space (W2CSpace) is proposed to improve the performance of pre-trained models . |
| Approach: | They propose a Word-Context-Coupled Space to replace pre-trained models with interpretable statistical logic. |
| Outcome: | The proposed language model can achieve better performance and highly credible interpretability compared to state-of-the-art methods. |
Three Real-World Datasets and Neural Computational Models for Classification Tasks in Patent Landscaping (2022.emnlp-main)
Copied to clipboard
| Challenge: | Patent Landscaping is one of the central tasks of intellectual property management and involves selecting and grouping patents according to user-defined technical or application-oriented criteria. |
| Approach: | They propose to use a novel model that takes into account textual information from the patents’ full texts as well as embeddings created based on the patent’s CPC labels. |
| Outcome: | The proposed model takes into account textual information from the patents’ full texts as well as embeddings created based on the patent’s CPC labels. |
A new European Portuguese corpus for the study of Psychosis through speech analysis (2022.lrec-1)
Copied to clipboard
| Challenge: | Psychosis is a clinical syndrome characterized by symptoms such as hallucinations, delusions, thought disorders and disorganized speech. |
| Approach: | They describe the creation of the first European Portuguese corpus for the identification of the presence of speech characteristics of psychosis. |
| Outcome: | The results show that spontaneous speech presents more identifiable characteristics than read speech to differentiate healthy and patients diagnosed with psychosis. |